Theoretical and Applied Genetics
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Theoretical and Applied Genetics's content profile, based on 49 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
de Freitas, G. M.; Certuche, D. S.; Jannink, J.-L.; de Oliveira, E. J.; Garcia, A. A. F.
Show abstract
Multi-trait genomic prediction offers a practical route to improve selection for costly, complex traits in clonally propagated crops such as cassava. In a Brazilian breeding panel of 1,078 cassava clones genotyped with 25,923 SNPs and phenotyped for six agronomic traits, we compared single-trait (ST) and multi-trait (MT) GBLUP models. Stage-wise mixed models produced BLUEs that fed into ST and MT-GBLUP. We tested five cross-validation schemes that mimic breeder realities: ST baseline (CV1); naive all-traits MT prediction for unphenotyped candidates (CV2); MT prediction using auxiliary trait phenotypes in the test set (CV3); and two sparse-phenotyping regimes with missingness by trait (CV4) or by clone (CV5) at 25%, 50%, and 75% levels. The main results were that, under the ST baseline (CV1), predictive ability ranged from 0.50 for DMC and 0.45 for FRY down to 0.13 for Le.Dis. A naive full MT model (CV2) performed approximately on par with ST-GBLUP. In contrast, MT designs (CV3) that included informative auxiliary traits, such as shoot yield and combinations with plant vigor and leaf disease severity, yielded small gains for DMC with predictive ability of approximately 0.51 (+2%), while FRY predictive ability increased to approximately 0.65 (+44%), accompanied by RMSE reductions for FRY up to approximately 13.5% (e.g. RMSE approximately 6.2). Sparse-phenotyping simulations (CV4/CV5) demonstrated that MT models sustain or even improve predictive ability under realistic missing-data regimes (PA {approx} 0.62 - 0.65). Selection concordance between MT and ST top-10% sets was generally high (>0.80), and MT configurations produced measurable improvements in expected selection response and genetic gain per cycle for several target traits. These results indicate that strategically implemented MT-GBLUP, using a small set of biologically and operationally informative auxiliary traits and optimized sparse phenotyping, can materially increase predictive accuracy and selection efciency for economically critical cassava traits while reducing phenotyping burden.
Kitony, J. K.; Reyes, V. P.; Sunohara, H.; Tasaki, M.; Yamasaki, M.; Mori, J.-i.; Shimazu, A.; Nishiuchi, S.; Michael, T. P.; Doi, K.
Show abstract
Genomic selection (GS) can accelerate genetic gain in crops, but its effectiveness depends on training population design and marker density. Nested association mapping (NAM) populations provide a structured framework that captures broad allelic diversity within a controlled genetic background. Here, we evaluated genomic prediction (GP) and genome-wide association study (GWAS) performance in an expanded aus-NAM population of rice comprising 1,818 recombinant inbred lines across 14 families and 11 agronomic traits, using genotyping-by-sequencing (GBS) markers and projected whole-genome sequence variants. Prediction accuracy plateaued at moderate marker densities ([~]20k SNPs) and with training populations of [~]500 lines ([~]40-60% of the available pool), with trait heritability emerging as the strongest determinant of predictive performance rather than model choice or marker density. In contrast, GWAS resolution continued to improve with increasing marker density, enabling detection of additional loci, including a chromosome 12 locus associated with heading date, while consistently recovering well-characterized genes such as EARLY HEADING DATE 1 (Ehd1) and SEMIDWARF 1 (SD1). These contrasting patterns indicate that GP reaches near-optimal performance once genome-wide variation is adequately represented, whereas GWAS benefits from higher marker density through improved locus resolution. The present study establishes a benchmark for implementing breeding programs involving japonica/indica crosses using GP in a single environment.
Salas, N.; Montazeaud, G.; Bourke, P. M.; Baranger, A.; David, J.
Show abstract
Modern agriculture faces major sustainability challenges, including stagnating yields, dependence on fossil resources, and severe environmental impacts. Increasing intra- and interspecific diversity within plots through agroecological design is a promising method for enhancing crop productivity and stability. However, mixed-crop performance remains highly variable, and the genetic architecture of interactions within heterogeneous canopies is poorly understood. Two quantitative genetic frameworks have been proposed: trait-based models, which describe how interacting traits shape phenotypes, and variance-based models, which treat neighbor genotype effects as "black-box" social effects. However, existing variance-based models have been developed almost exclusively for intraspecific interactions and simple neighborhoods. We propose a general multispecies framework describing how a focal plants phenotype and total breeding value arise from its own direct effects and from the indirect effects of conspecific and heterospecific neighbors. We derived analytical expressions for phenotypic variance, inter-individual covariance, total breeding value variance, and relative heritable variance, which explicitly account for spatial structure, relatedness, and environmental similarities. Using a two-species alternating-row field layout and extensive simulations based on flexible variance-covariance structures, we evaluated the statistical power and bias of joint mixed-model estimators of direct and indirect genetic and environmental effects under a wide range of parameter combinations. Our results show that accurate separation of direct and indirect effects depends on trait heritability and replication, and that modeling genetic covariances across effects and species substantially improves estimation accuracy. This framework provides a unified, individual-centered basis for analyzing complex multispecies neighborhoods and quantifying the breeding potential of plant communities. Article SummaryGrowing several crop species or varieties together in the same field can boost yield and stability, but the outcome is unpredictable and the genetic causes remain unclear. We developed a theoritical & statistical framework that links each plants performance to its own genes and to those of its neighbors, both from the same and from a different species. Computer simulations of a two-species field showed that these direct and neighbor-driven genetic effects can be reliably separated when enough plants are measured per variety. The framework opens the way to breeding crop mixtures that perform well specifically when grown alongside another species.
Aldiss, Z.; Brunner, S.; Heidariask, B.; Chenu, K.; Van Haeften, S.; Baraibar, S.; Ganesgalingam, D.; Moody, D.; Hickey, L.; Lam, Y.
Show abstract
PurposeGenotype-by-environment (G x E) interactions represent a major obstacle to increasing genetic gain in crop breeding, with the underlying physiological drivers often remaining obscured within conventional statistical models. This case study presents a novel framework that transforms the latent factors from Factor Analytic (FA) multi-environment trial (MET) models into heritable quantitative traits, enabling the genetic dissection of adaptive response patterns. MethodsA Factor Analytical Linear Mixed Model (FA-LMM) was fit to plot-level yield data for 1,036 barley genotypes across eight Australian trials. ResultsCorrelation of the factor loadings with APSIM-simulated environmental covariates demonstrated that the second latent factor FA2 was strongly correlated with the Water Stress Index (r = -0.83) during the critical flowering period, establishing water availability as the main biological axis of crossover Gx E. Genotypic scores for the derived traits, Overall Performance (OP) and Water Stress Response (WSR), were subjected to high-resolution haplotype-based mapping using local Genomic Estimated Breeding Values (GEBV). ConclusionThis analysis successfully identified major genomic regions that accounted for a substantial proportion of the additive genetic variance. Gene Ontology enrichment of candidate genes within the top haploblocks implicated fundamental pathways related to energy homeostasis, root development, and stress response, with notable candidates including FTsH11, BPS1, and TDP1. The distribution of favourable Haplotypes of Interest (HOI) in elite cultivars suggested a historical signature of inadvertent selection for these adaptive mechanisms. This framework provides an explicit bridge between statistical modelling and functional genomics, offering breeders actionable genetic targets for accelerated development of climate-resilient cereals.
Sakurai, K.; Moreau, L.; Mary-Huard, T.; Charcosset, A.; Iwata, H.
Show abstract
In plant breeding, it is often necessary to improve a target trait while maintaining other essential traits within desirable ranges. When genetic relationships exist among these traits, improvements in the target trait may lead to undesirable changes in essential traits, complicating cross selections. In such cases, it is critical to select cross-pairs that are expected to produce progeny that satisfy the requirements for all traits. The progeny distribution of each crossing pair can be predicted using the estimated genotypic values and genetic (co)variances of the target and essential traits. By utilizing this distribution, the probability of generating progeny that satisfy predefined trait requirements can be evaluated, allowing a direct comparison of alternative crosses. In this study, we developed Cross Potential Selection for Multiple Traits (CPS-MT), a breeding strategy designed to improve a target trait while maintaining one or more essential traits within desirable ranges. CPS-MT extends the original Cross Potential Selection (CPS) framework to explicitly handle trade-offs between traits under genetic correlations. We evaluated the performance of CPS-MT through simulations involving four types of genetic relationships and two genetic causal factors between traits, resulting in seven scenarios. Across all scenarios, CPS-MT consistently improved the likelihood of obtaining desirable progeny, indicating that CPS-MT provides a practical and effective framework for cross selection under multi-trait constraints in breeding programs. Article SummaryThis study developed Cross Potential Selection for Multiple Traits (CPS-MT), a new breeding strategy designed to improve a target trait while maintaining one or more essential traits within desirable ranges. CPS-MT evaluates crossing pairs by predicting progeny distributions based on estimated genotypic values and genetic covariances, enabling direct comparison of alternative crosses under multi-trait constraints. Through simulations incorporating four types of genetic relationships and two causal factors (seven scenarios), CPS-MT consistently increased the likelihood of obtaining progeny that satisfied the predefined trait requirement. These results indicate that CPS-MT provides a practical, robust framework for target trait improvement under trait constraints.
Li, Z.; Li, X.; Liu, S.; Wilson, I.; Zhu, Q.-H.; Stiller, W.; Conaty, W.
Show abstract
Genomic prediction (GP) across diverse environments has a potential to accelerate genetic gain in cotton breeding programs. A major challenge in GP is modelling genotype-by-environment interactions (GEI), which is essential for selecting stable and high-performing genotypes under variable production conditions. However, incorporating GEI into GP models increases the dimensionality and computational complexity, risking complex models that are impractical to use on commercial breeding-scale data sets because of run times and computational demands. This study addresses two primary aims. Firstly, we evaluate the practical benefits of GEI-informed GP for predicting economically important cotton traits. Second, advanced statistical modelling strategies are developed and assessed for integrating genomic and environmental data at scale. We propose a dimensionality reduction approach that combines linkage disequilibrium network analysis with principal component techniques to reduce redundancy while preserving informative variation. Using this reduced dataset, we implement Bayesian linear regression models and, for comparison, deep residual neural networks for genomic prediction. Analyses were conducted on a large multi-environment dataset from the CSIRO cotton breeding program, comprising 3,236 breeding lines, 54 environmental covariates, and 8,049 yield and fibre quality phenotype records collected over 10 years and 9 locations representing 41 year-location combinations. Results demonstrate that generally Bayesian linear regression approaches outperform BG-BLUP models, with all three linear/linear mixed methods providing clearly more reliable performance than the deep learning models. These findings highlight the value of using interpretable statistical models for integrating genomic and environmental information to support selection decisions under diverse environmental conditions.
Montesinos-Lopez, O. A.; Montesinos-Lopez, A.; Montesinos-Lopez, J. C.; Crossa, J.; Dreisigacker, S.; Hernandez-Suarez, C. M.; Ortiz, R.
Show abstract
Accurate modeling of genotype-by-environment (GxE) interaction is critical for genomic prediction in plant breeding but remains challenging due to complex interaction structures. Conventional models often use the Hadamard product of genotype and environment covariance matrices to capture joint similarity, which may not fully represent GxE complexity. Here we propose a novel framework that derives covariance structures from the matrix multiplication of genotype and environment kernels, decomposing these into symmetric components incorporated as random effects in mixed models. Evaluated for 11 wheat and rice multi-environment datasets and across, this approach consistently outperformed the traditional Hadamard-based model, improving prediction accuracy by up to 13.2% in Pearsons correlation and enhancing top-selection accuracy. Combining both methods yielded the highest performance, indicating complementary information capture. This framework offers a flexible, interpretable, and computationally feasible extension for modeling GxE interaction, potentially enhancing genomic selection effectiveness under diverse environmental conditions.
Hamazaki, K.; Tsuda, K.
Show abstract
Background: Germplasm collections contain wide genetic diversity that is valuable for plant breeding, but conducting phenotypic evaluation for all genotypes in field trials is rarely feasible. Bayesian optimization offers a way to decide, season by season, which genotypes to cultivate in order to identify superior genotypes with fewer evaluations. However, standard Bayesian optimization commonly starts from randomly selected genotypes and mainly relies on surrogate models built from marker genotype information, while the text-based passport information that accompanies germplasm is not fully used. We examined whether pre-trained large language models can provide prior knowledge that improves these decisions in germplasm evaluation. Results: We constructed a large-language-model-guided Bayesian optimization framework that introduces large language models into two parts of the Bayesian optimization workflow. In zero-shot warmstarting, a large language model proposes initial genotypes using passport information such as cultivar name, country of origin, and subpopulation, optionally together with principal component scores derived from genome-wide single-nucleotide-polymorphism markers. In addition, we evaluated a large-language-model-based surrogate model that predicts phenotypic values for untested genotypes using in-context learning from previously evaluated genotypes. Using a rice germplasm panel and two target traits (seed number per panicle for maximization and protein content for minimization), we compared strategies. For seed number per panicle, zero-shot warmstarting with a general-purpose instruction-following model reduced the number of evaluated genotypes needed to reach the best genotype, whereas improvements were small for protein content. When genomic information was available, Gaussian-process-based Bayesian optimization was the strongest overall approach, while the large-language-model-based surrogate model outperformed random baselines and was competitive in some settings. When genomic information was not available, predictions based on passport information improved efficiency compared with fully random strategies. Conclusions: Pre-trained large language models can inject useful agronomic knowledge into Bayesian optimization for germplasm evaluation, particularly by improving early-stage genotype selection, and can also support optimization when genomic information is unavailable. As models better handle long genomic sequences together with passport information, large-language-model-guided Bayesian optimization may become a practical and explainable decision-support approach for agricultural optimization.
Kinoshita, S.; Iwata, H.
Show abstract
Intercropping is a promising strategy to improve productivity and sustainability in agricultural systems, but designing effective genotype combinations remains a major challenge owing to the rapid increase in possible pairings as the number of candidate genotypes increases. This creates a practical bottleneck because field evaluation of all combinations is infeasible under realistic resource constraints. Here, we propose a framework that integrates genomic prediction and Bayesian optimization to support efficient decision-making for intercropping system design. Using genome-wide marker data from sorghum and soybean, we simulated intercropping performance across 5,214 genotype pairs under certain genetic architectures, including variation in heritability, correlations between direct and indirect genetic effects, and the contribution of pair-specific interactions. Genomic prediction models incorporating direct and indirect genetic effects substantially improved prediction accuracy compared with models based on direct genetic effects alone, and inclusion of specific mixing ability further enhanced the performance under high-heritability conditions. When coupled with Bayesian optimization, the models rapidly identified superior genotype pairs, requiring fewer evaluation cycles than random or prediction-only search strategies. Acquisition functions that account for predicted uncertainty were most effective in complex scenarios involving interaction effects or negative correlations between direct and indirect effects. These results demonstrate that combining genomic prediction with Bayesian optimization can substantially reduce the experimental burden associated with intercropping design, while improving the efficiency of identifying high-performing genotype pairs. The proposed framework provides a practical approach for prioritizing candidate mixtures in breeding and field evaluation, and contributes to the development of data-driven strategies for sustainable agricultural systems. HighlightsO_LIA data-driven framework was developed to optimize genotype pairs in intercropping. C_LIO_LIModeling indirect effects improved prediction accuracy across genotype pairs. C_LIO_LIPair-specific interactions enhanced prediction under high-heritability conditions. C_LIO_LIBayesian optimization identified superior pairs under limited evaluation capacity. C_LIO_LIThe framework reduces field-testing requirements for intercropping system design. C_LI
Ingold, M.; Gao, Q.; Mandel, J. R.; McNellie, J. P.; Keepers, K. G.; Barb, J. G.; Burke, J. M.; Rieseberg, L. H.; Hulke, B. S.
Show abstract
In sunflower (Helianthus annuus L.), the composition of fatty acids in the seeds, primarily oleic, linoleic, stearic and palmitic acid, is of utmost importance for oil quality. Despite this, the genetic basis of this trait and its interaction with the environment is poorly understood. Understanding this interaction is critical to improvement of sunflower within the context of climate change. In this work, we incorporated fatty acid composition measurements from the sunflower SAM population and eight environments across an extensive geographic cline into GWAS. The SAM panel consists of 287 varieties representing approximately 90% of sunflower diversity, for which 2.2 million high-quality SNPs with a MAF > 5% are available. For increased power, multivariate GWAS was performed with four different inputs: (i) mean fatty acid composition within each environment, (ii) mean fatty acid composition within each environment omitting high oleic varieties, (iii) trait stability within environments quantified by standard errors among replicate samples ( stability) and (iv) Eberhart and Russells {beta} which quantifies trait stabilities across environments ({beta} stability). All four analyses yielded highly significantly associated SNPs. We found that high oleic varieties exhibited high {beta} trait stability, resulting in substantial overlap in markers between analyses (i) and (iv), with signals being fairly consistent between environments in analysis (i). For analyses (ii) and (iii), significant markers tended to vary between trials. For significant SNPs across all analyses, 147 candidate genes were identified, including promising candidates such as 15 fatty acid metabolism genes, 6 heat shock proteins and 22 transcription factors. Lastly, a large introgression consisting of two flanking inverted sequences on Chromosome 5 was found to coincide with stability in the Georgia trial, suggesting a role in FA composition stability under high heat conditions.
Acharya, S. R.; Garcia-Abadillo, J.; Lyerly, J.; Brown-Guedira, G.; Jarquin, D.; Bandillo, N.
Show abstract
Genomic prediction models that account genotype-by-environment (GxE) have the potential to accelerate the rate of genetic gain for yield and agronomic performance, yet relatively few studies have applied GxE prediction in public soft red winter wheat (Triticum aestivum) breeding programs. In this study, we extended a reaction norm-based genomic prediction framework by integrating weather-based environmental covariates to more effectively capture genotype- environment interactions. Key agronomic traits, including seed yield, plant height, test weight, and heading date, were evaluated across 33 environments (location-year) using over 3,200 breeding lines from the North Carolina State University small grains breeding program. Multiple genomic prediction models were compared using several cross-validation (CV) schemes representing common breeding scenarios. Across traits, the reaction norm M5 model, which incorporates both GxE and genotype-by-environmental covariate interactions (GxO), achieved the highest prediction accuracy (PA) in CV2 (predicting incomplete field trials) and CV1 for yield and test weight (predicting new lines). The highest PA was observed for test weight under CV2 (0.54) and for yield under CV1 (0.41). Under CV0 (predicting new environments), the M3 model incorporating GxE produced highest PA across traits, with the greatest accuracy for plant height (0.45), although differences among M2, M3, and M4 were small. Prediction under CV00 (predicting new lines in new environments) remained more challenging, with PA values 0.10 - 0.20 across traits. Overall, our results demonstrate that integrating environmental covariates into genomic prediction models can improve predictive performance across diverse wheat-growing environments in North Carolina, supporting their utility for applied breeding efforts. CORE IDEASO_LIIntegrating genotype-by-environment (GxE) interactions with environmental covariates improves prediction accuracy across environments. C_LIO_LIModel performance varies by prediction scenario, with different approaches performing best for new lines, incomplete trials, or new environments. C_LIO_LIPrediction of new lines in new environments remains challenging. C_LI PLAIN LANGUAGE SUMMARYThis study explores how adding environmental information to genomic prediction models can improve prediction accuracy in a public winter wheat breeding program. Using data from multi-environment trials conducted across diverse conditions in North Carolina, we evaluated statistical models that capture how different wheat lines respond to changing environments. By incorporating weather data, we improved the ability to predict performance across locations and years. These findings provide practical insights for refining selection strategies and accelerating genetic gain in wheat breeding.
Shaffer, W.; Papin, V.; Carter, Z.; Brunner, S. M.; Tong, J.; Villiers, K.; Robinson, H.; Voss-Fels, K.; Hayes, B. J.; Hickey, L.; Dinglasan, E.
Show abstract
Haplotype-based breeding strategies have emerged as promising approaches to maximize long-term genetic gain by identifying complementary parental combinations while maintaining genetic diversity. However, these methods typically require phased genotypes and more intensive workflow pipelines and skillsets. We developed a novel local genomic estimated breeding value (localGEBV) fitness function with similar intent to the optimal haplotype stacking (OHS) framework fitness function and implemented both in the novel R package, HapSelect. Our aim was to evaluate whether phased haplotypes provide additional benefit over the more easily available dosage-based unphased genotypes in highly inbred crops. A subset of bread wheat nested association mapping (NAM) population comprising 444 lines genotyped with 6,054 DArT-Seq markers was analysed. Marker effects were estimated using rrBLUP, localGEBV and haplotype effects were calculated across linkage disequilibrium-defined haploblocks, and genetic algorithms (GA) were used to identify optimal sets of 30 founders using either a localGEBV derived fitness function with unphased, dosage inputs or the OHS fitness function with phased inputs. Selected parental sets were compared with conventional truncation selection (TS) through 150 generations of forward simulation. The OHS fitness function achieved a marginally greater optimized ultimate GEBV than the localGEBV fitness function during GA optimization, with only 18 of the 30 selected founders overlapped between the two methods. Despite these differences, forward simulations demonstrated nearly identical long-term genetic gain for localGEBV and OHS-selected founders, with both approaches outperforming conventional truncation selection by maintaining greater genetic diversity and delaying the genetic plateau. The minimal difference between localGEBV and OHS is likely attributable to the high homozygosity of the population, where localGEBV and haplotype effects are nearly confounded. These results demonstrate that dosage-based localGEBV provides a practical alternative to phased haplotype approaches for parent selection in inbred crops, substantially simplifying genomic workflows while maintaining long-term breeding performance. Future work should evaluate these methods in more diverse inbred populations and outbred species, where great haplotypic diversity may increase the advantage of true haplotype-based optimizations.
Santos Junior, D. R. d.; Fe, D.; Lenk, I.; Jensen, C. S.; Asp, T.; Janss, L.; Bornhofen, E.
Show abstract
The performance of a single cross is determined by the average additive effects of the parents, as well as the interactions between them. These quantities can be estimated using an appropriate genetic design, allowing for the estimation of general (GCA) and specific (SCA) combining abilities. The prediction of GCA for new parents and the total genetic value of unrealized crosses can be made when genome-wide marker information is available. Several studies in crops such as maize and rice have demonstrated the potential of genomic-assisted prediction of single-cross performance in economically important crops. However, no study to date has explored its relevance in perennial ryegrass, an obligate allogamous species that is bred in genetically heterogeneous families. In this study, we aimed to estimate genetic parameters and assess the ability of genomic models to predict the performance of F2 families in terms of dry matter yield and nutritive quality traits. We used data from a large partial diallel involving 104 parents from two distinct subpopulations, as inferred by admixture analysis. F2 families were evaluated in multiple environments and under two nitrogen availability conditions. Genotyping-by-sequencing of the parent plants produced 42,145 variants after quality control, which were used to estimate genomic relationships based on identity-by-state. Variance component estimation revealed limited GCA and SCA interactions with the environment, and particularly with nitrogen management. The predictive abilities of two parental models exceeded 0.60 and often surpassed 0.70 for most traits. However, incorporating non-additive effects into the model did not improve predictive ability. We leveraged the genetic diversity among parents to map genomic regions associated with all recorded traits. Genome-wide association studies (GWAS) by genomic best linear unbiased prediction (GBLUP) identified six quantitative trait loci (QTL) regions, with 45 candidate genes within the linkage disequilibrium range, estimated at approximately 92 kb. Our results demonstrate that genomic prediction of single crosses can be performed with high accuracy, especially when both parents are also progenitors of families in the training set.
Johansen, N. H.; Sarup, P.; Hansen, P.; Orabi, J.; Jahoor, A.; Ramstein, G. P.
Show abstract
In quantitative genetics, candidate SNPs are identified through genotype-phenotype associations inferred with genome-wide association studies (GWAS). In this study, we explore an alternative approach to detect genetic variants with non-neutral effects by tracking temporal trends in allele frequency in a winter wheat (Triticum aestivum L.) breeding population over an eight-year period, from which signals of selection may be inferred. Selection signatures were inferred with a generalized linear model, where we modeled trends in allele frequency as a function of time (crossing year). These signatures of selection were used to prioritize variants. Associations between phenotypic performance and individual load of prioritized variants were then investigated. Furthermore, we assessed whether incorporating selection information into a genomic best linear unbiased prediction (GBLUP) model improves model performance in terms of quality of fit and prediction ability. Our findings indicate that the inferred signals of selection are effective in identifying non-neutral variants. Variants under strong negative selection were associated with a decrease in protein content adjusted for grain yield (p-value < 0.01), while genetic variants that had been under moderate to high levels of positive selection were associated with increased grain yield (p-value < 0.01). However, incorporating selection information did not improve prediction accuracy. In conclusion, temporal trends in allele frequency can be used to detect non-neutral variants. The proposed approach may hence complement traditional quantitative genetic methods for detecting non-neutral genetic variation. This approach may allow breeders to detect non-neutral variants earlier in the breeding cycle, without resorting to phenotypic data.
El Ghazzal, Z.; Pegard, M.; Guacaneme, M.; Surault, F.; Arcia-Ruiz, I.; Julier, B.
Show abstract
Lucerne is gaining interest as a living mulch in agroecological productions. However, its vigorous growth can lead to competition with cash crops for light and nutrients, necessitating new ideotypes. This study investigated the genetic basis of traits relevant to ideotype breeding: dormancy, spring regrowth, height, growth habit, leaflet size, stem diameter, and plant structure. Individuals from a diversity panel of 27 accessions and a synthetic population were phenotyped in a spaced plant nursery. Over 100,000 SNP markers were used for genotyping. Genome-wide association study (GWAS) and genomic prediction were conducted, considering population structure. Heritability estimates ranged from moderate to high in diversity panel (h{superscript 2} = 0.36-0.70) but were lower in synthetic population (h{superscript 2} = 0.17-0.33), reflecting reduced genetic variance. Trait correlations differed markedly between populations, indicating the possibility of recombining traits to create new ideotypes. GWAS identified a few QTL (r{superscript 2} up to 0.27) for leaflet size, height, growth habit, and plant structure, with candidate genes linked to growth, stress response, and signalling pathways. Genomic prediction was highly accurate in diversity panel, where broad genetic variation allowed reliable estimation of marker effects, with prediction accuracies exceeding 0.8 for heritable traits, including growth habit and leaflet size. In contrast, accuracies were low in synthetic population, reflecting its limited diversity and small size, whether training was based on the synthetic population itself or on the diversity panel. These results highlight the potential to recombine traits and develop lucerne ideotypes using molecular tools such as QTL detection and genomic prediction.
Tajima, A. M.; Matthews, W. C.; Duong, T.; Khanh, T. D.; Baniya, A.; Penmetsa, R. V.; Parker, T.; Farmer, A.; English, S.; Diepenbrock, C.; Gepts, P.; Roberts, P. A.; Huynh, B.-L.
Show abstract
Lima bean (Phaseolus lunatus) is a broadly adapted, economically important leguminous crop and a susceptible host of root-knot nematodes (Meloidogyne spp.; RKN), which are a devastating plant pathogen in agricultural systems worldwide. To date, there have been few studies to elucidate the genetic determinants of RKN resistance in lima beans. Understanding the genetic mechanisms underlying resistance is essential for improving resistance traits and incorporating them into lima bean breeding programs. To assist in marker-assisted selection, we aimed to identify and map quantitative trait loci (QTLs) conferring RKN resistance-related traits. Three recombinant inbred line (RIL) populations were used in this study. Three populations were derived by crossing two RKN-resistant parents with the same RKN-susceptible parent and with each other. All populations were genotyped using genome-wide single-nucleotide polymorphism (SNP) markers. Each population was screened for root galling (RG) and RKN egg reproduction (ER) in response to M. incognita and M. javanica in greenhouse experiments. Three major QTLs were detected and mapped on chromosome Pl04 (QRk-pl04.1), Pl05 (QRk-pl05.1) and Pl10 (QRk-pl10.1) across populations. Among them, QRk-pl05.1 and QRk-pl10.1 affected levels of RG and ER of both RKN species, while QRk-pl04.1 suppressed root galling and reproduction responses of M. incognita but not of M. javanica. These chromosomal regions defined by flanking markers will help guide marker-assisted breeding and gene discovery for broad-based RKN resistance in lima beans.
Parthasarathy, S.; Rocheford, T.; Koehler, K.
Show abstract
BackgroundDecades of maize (Zea mays L.) QTL mapping have produced fragmented results across hundreds of independent studies, characterized by broad confidence intervals, population-specific effects, and a predominantly single-trait analytical scope. Comprehensive multi-trait integration remains limited, yet it could substantially improve our understanding of trait relationships for strategic breeding. We integrated 2,701 QTLs published over 30 years across five functionally distinct trait categories (grain yield and components; plant development and architecture; plant physiology and stress adaptation; grain quality and nutritional composition; and disease and pest resistance) in order to identify functionally classified genomic hotspots and prioritize candidate genes for multi-trait breeding applications. ResultsBioMercator V4.2 consolidated 2,518 projectable QTLs into 187 high-confidence meta-QTLs (MQTLs), achieving an average 59% reduction in confidence interval width; 128 of 187 MQTLs (68.4%) achieved dual-platform support through GWAS co-localization. Twenty-three genomic hotspots harbored 132 of 187 MQTLs (70.6%) and were classified into three functional categories: twelve multi-trait hubs that may enable simultaneous improvement of multiple traits through pleiotropic or tightly linked genes; seven single-trait clusters with pathway-specific effects, exemplified by the chromosome 9 starch biosynthesis cluster; and four major-effect loci with reported individual effects exceeding 20% PVE, including vgt1 (54% PVE) and opaque2 (34.2% PVE). Descriptive environmental classification distinguished MQTLs predominantly supported by optimal-condition QTLs (42%) from those predominantly supported by stress-condition QTLs (28%), the latter showing approximately 3.5-fold greater mean contributing-QTL phenotypic variance, directionally consistent with conditional genetic effect amplification under stress. Network-based candidate gene prioritization combined with cross-cereal ortholog analysis showed that 67% of the top candidates possess orthologs in rice, sorghum, wheat, or barley, and 53% are conserved across all four species, identifying priority targets for functional genomics investment. ConclusionsThis functionally classified and environmentally characterized meta-QTL framework provides breeders with a structured resource for multi-trait hotspot selection, environment-appropriate allele deployment, and functional genomics prioritization, with broader applicability as a transferable analytical template for other crop species confronting analogous challenges of fragmented QTL literature and complex multi-trait breeding objectives.
Solarte Certuche, D. C.; Mamedio de Freitas, G.; Jannink, J.-L.; Garcia Morales, C. F.; Sousa Cerqueira, T.; Santos de Santana, B.; Jorge de Oliveira, E.; Garcia, A. A. F.
Show abstract
Cassava is a major staple crop in tropical regions, and improving its root nutritional quality, particularly carotenoid and dry matter content (DMC), remains a central breeding goal. To elucidate the genetic basis of these traits by locating genomic regions associated with them, we analyzed 3,043 cassava clones from the Brazilian Agricultural Research Corporation (Embrapa) breeding program, phenotyped across 188 multi-environment trials conducted from 2011 to 2022 in Brazil. All clones were genotyped using Genotyping-by-Sequencing (27,045 Single Nucleotide Polymorphism - SNPs) and Diversity Arrays Technology - DArTseq (25,923 SNPs). Trait values were estimated using a two-stage mixed model to obtain deregressed BLUPs (Best Linear Unbiased Predictions), and genome-wide association analyses were performed using both the Mixed Linear Model (MLM) and Multi-Locus Mixed Model (MLMM). We detected six significant SNPs consistently associated with carotenoid content and DMC after Bonferroni correction. These SNPs mapped to six candidate genes involved in pathways relevant to root physiology, including Abscisic Acid ABA-related signaling, hydrolase activity affecting carotenoid conversion, fatty-acid biosynthesis within plastids, cell-wall remodeling, and glycolytic energy metabolism. The loci jointly explained 75.56 % of the phenotypic variance for carotenoids and 76.23 % for DMC, with individual SNP effects ranging from [~]17 % to [~]42 % PVE (Proportion of Variance Explained). Broad-sense heritability was H2 = 0.78 for carotenoids and H{superscript 2} = 0.34 for DMC, confirming substantial genetic control and suitability for molecular breeding. Haplotype analyses revealed four superior haplotypes for carotenoids and one key haplotype for DMC, each showing significantly higher trait values compared with other allelic combinations. These haplotypes represent promising targets for marker-assisted selection and genomic selection, with direct applicability for accelerating genetic gain in elite breeding populations. The results provide actionable genomic resources for breeding programs aiming to develop biofortified and high-root quality cultivars and establish a foundation for future multi-omics and functional validation studies.
Baraja-Fonseca, V.; Gil-Villar, D.; Bancic, J.; Renau-Morata, B.; Salud Justamante, M.; Plazas, M.; Gramazio, P.; Vilanova, S.; Perez-Perez, J. M.; Granell, A.; Molina, R. V.; Nebauer, S. G.; Prohens, J.; Arrones, A.
Show abstract
Nitrogen-use efficiency (NUE) is a pivotal breeding target in tomato (Solanum lycopersicum L.) to sustain production under reduced N inputs. Here, we leveraged a recently developed tomato multi-parent advanced generation inter-cross (ToMAGIC) population to identify lines with superior performance under reduced N availability. The eight founders and a core subset of 118 ToMAGIC lines were characterized with 10,684 SNP markers and evaluated under optimal (opN, 15 mM) and suboptimal (subN, 8 mM) N supply in an experiment totalling 1,576 plants, generating 48,068 data points across 61 phenotypic variables. Under both N treatments, ToMAGIC lines exhibited transgressive segregation for most traits, confirming the value of this population as a reservoir of untapped variation. Notably, under subN conditions, harvest index (Hi) increased by 29-44%, suggesting adaptive resource redistribution toward reproductive sinks. Variance partitioning revealed that agronomic and NUE-related traits were largely under genetic control, with heritability estimates frequently above 0.80 and broadly conserved across N treatments. Multivariate trait analysis identified fruit yield N concentration (NUE component, CN,y), shoot biomass N content (NAb), and shoot growth-related traits as the main drivers of treatment differentiation. Finally, proxy traits were prioritized by integrating response magnitude, heritability, trait correlations, and treatment-discriminatory power into multi-trait selection indices. This strategy generated favorable predicted genetic gains, reaching 158% for high-performance lines and 170% for subN-adapted lines, and consistently identified lines 402, 428, 518, 800, and 816 as promising pre-breeding materials. Overall, this study supports ToMAGIC as a powerful resource for developing N-efficient cultivars suited for sustainable agriculture.
Ara, A. M.; Holmes, D. J.; Friesen, T. L.; Carver, B. F.; Bai, G.; St. Amand, P.; Bernado, A.; Sharma, R.; Aoun, M.
Show abstract
Key message Characterized and unknown septoria nodorum blotch susceptibility/resistance genes were identified in contemporary U.S. hard winter wheat. The necrotrophic fungus Parastagonospora nodorum is the causal agent of septoria nodorum blotch (SNB) of wheat. To determine the prevalence of SNB sensitivity genes in a contemporary U.S. hard winter wheat (HWW), we evaluated a panel of 619 breeding lines and cultivars against five P. nodorum isolates and five necrotrophic effectors (NEs), SnToxA, SnTox1, SnTox3, SnTox267 and SnTox5, and genotyped the panel using genotyping-by-sequencing (GBS) markers and diagnostic Kompetetive-allele specific PCR (KASP) markers for the sensitivity genes Tsn1-B1, Snn1-B1, and Snn3-B1/B2. GBS analysis identified 34,357 GBS-single nucleotide polymorphism (SNP) markers. Evaluations against P. nodorum isolates showed that 40-67% of the genotypes were susceptible in the panel. Toxin infiltration assays showed that 54%, 2%, 37%, 13%, and 15% of the genotypes were sensitive to SnToxA, SnTox1, SnTox3, SnTox267, and SnTox5, respectively. Diagnostic KASP markers for Tsn1-B1, Snn1-B1, and Snn3-B1/B2 showed prediction accuracies of 98%, 75%, and 92% for the corresponding effectors SnToxA, SnTox1, and SnTox3, respectively. Genome-wide association studies (GWAS) not only confirmed the presence of the previously characterized sensitivity genes Tsn1-B1, Snn1-B1, Snn2, Snn3-B1/B2, and Snn5-B1, but also identified new loci to be associated with responses to P. nodorum isolates and NEs. Of which, Qsnb.osu-2AS on chromosome 2AS was associated with responses to all five isolates. We developed KASP markers KASP_S4B_643615365, KASP_ S2D_16184991, and KASP_S2A_9833162 linked to Snn5-B1, Snn2, and Qsnb.osu-2AS, respectively. These findings should guide breeding for SNB resistance in hard winter wheat.